Skip to content

fix(bin): sync upstream reliability and compatibility fixes - #97

Merged
yelenplays merged 24 commits into
mainfrom
fm/firstmate-upstream-sync-v4
Oct 3, 2026
Merged

yelenplays merged 24 commits into
mainfrom
fm/firstmate-upstream-sync-v4

Conversation

@yelenplays

@yelenplays yelenplays commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Upstream port point: 1f3e769

Intent

Update everything: bring Firstmate up to date with upstream ("I want to update everything").

Context: this repository is a fork (origin yelenplays/firstmate) of upstream kunchenguid/firstmate. Upstream main has 20 commits the fork does not yet contain, including Pi 1.0 compatibility (kunchenguid#6338), Pi watcher continuity (kunchenguid#5489), titled Claude top rules in the composer classifier (kunchenguid#5963), and polling-churn, watcher-reclaim and ShellCheck-memory fixes. The fork's own behavior, including its fork-only commits, must survive the update.

What Changed

  • Merge upstream main’s 20 commits while preserving fork-specific changes.
  • Improve watcher continuity across Pi arm gaps and supervision-host hand-backs; add Pi compatibility and seeded-home trust handling.
  • Reduce remote-job polling churn and harden captain-hold validation, Claude composer classification, and ShellCheck memory fallback.

Risk Assessment

⚠️ Medium: The sync brings in broad changes across watcher continuity, remote jobs, inbox handling, and linting; no concrete defect was found, but the cross-cutting scope warrants caution.

Testing

Exercised targeted Pi watcher and extension tests, composer classification, remote-job polling, watcher recovery, and ShellCheck memory fallback; the real installed Pi SDK passed continuity and branch handoff checks. The supervision-host suite hit its time bound after the relevant reclaim case passed. Pi 1.0 was unavailable, and tmux was not on PATH, so no Pi 1.0 or interactive TUI run or screenshot was possible.

  • Live validation: ⚠️ inconclusive - 1 of 6 scenarios driven live against the product
Scenario Result Live Evidence
A Pi session processes watcher wakes across successor handoff without losing queued work. ✅ pass live tests/fm-pi-branch-live-e2e.test.sh against installed Pi SDK 0.87.1
A user runs the updated extensions on Pi 1.0 and Pi accepts their runtime/session behavior. ⏸️ untested no Pi 1.0 is not installed. Provide a Pi 1.0 runtime/SDK in the test environment and rerun the compatibility scenario.
An idle Claude composer with a titled top rule is recognized as empty, while a draft remains pending. ⏸️ untested no The prior payload only records tests/fm-composer-lib.test.sh, not a live composer interaction. Provide a running product with the composer available and drive both the empty and draft states.
Queued remote work preempts polling without losing its cursor or causing sibling-poll churn. ⏸️ untested no The prior payload only records tests/fm-remote-job.test.sh, not live remote work in the product. Provide a running product with queued remote work and drive the polling/preemption scenario.
After a pass-through leaves a watcher cycle orphaned, the next park reclaims it so one arm owns the cycle. ⏸️ untested no The prior payload records a passing test case in tests/fm-supervision-host.test.sh, but does not establish that it was driven against the live product. Provide a running product and drive the pass-t…
When ShellCheck hits a memory failure, lint retries without external sources and preserves genuine fallback findings. ⏸️ untested no The prior payload only records the memory-failure fallback test in tests/fm-lint.test.sh, not a live product lint run. Provide a running product and drive the ShellCheck memory-failure fallback scen…
Evidence: Live Pi SDK continuity and handoff
ok - real Pi SDK 0.87.1 queues a streaming-time watcher wake without before_agent_start, keeps the successor chain, and surfaces consumption of both follow-ups
Evidence: Watcher reclaim case
ok - host+hook: the next park takes over the cycle a main-only pass-through left, so one arm owns it
- Outcome: ⚠️ 2 warnings across 1 run (32m56s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

⚠️ **Test** - 2 warnings
  • ⚠️ The required Pi 1.0 compatibility scenario could not be exercised: the installed Pi CLI and SDK are both 0.87.1. The supervision-host test also exceeded its 1,200-second bound after the specific next-park watcher-reclaim check passed, leaving later cases unverified.
  • ⚠️ live validation verdict: inconclusive (1 of 6 scenarios were driven live against the product); untested: A user runs the updated extensions on Pi 1.0 and Pi accepts their runtime/session behavior., An idle Claude composer with a titled top rule is recognized as empty, while a draft remains pending., Queued remote work preempts polling without losing its cursor or causing sibling-poll churn., After a pass-through leaves a watcher cycle orphaned, the next park reclaims it so one arm owns the cycle., When ShellCheck hits a memory failure, lint retries without external sources and preserves genuine fallback findings.
  • Live validation: ⚠️ inconclusive - 1 of 6 scenarios driven live against the product
Scenario Result Live Evidence
A Pi session processes watcher wakes across successor handoff without losing queued work. ✅ pass live tests/fm-pi-branch-live-e2e.test.sh against installed Pi SDK 0.87.1
A user runs the updated extensions on Pi 1.0 and Pi accepts their runtime/session behavior. ⏸️ untested no Pi 1.0 is not installed. Provide a Pi 1.0 runtime/SDK in the test environment and rerun the compatibility scenario.
An idle Claude composer with a titled top rule is recognized as empty, while a draft remains pending. ⏸️ untested no The prior payload only records tests/fm-composer-lib.test.sh, not a live composer interaction. Provide a running product with the composer available and drive both the empty and draft states.
Queued remote work preempts polling without losing its cursor or causing sibling-poll churn. ⏸️ untested no The prior payload only records tests/fm-remote-job.test.sh, not live remote work in the product. Provide a running product with queued remote work and drive the polling/preemption scenario.
After a pass-through leaves a watcher cycle orphaned, the next park reclaims it so one arm owns the cycle. ⏸️ untested no The prior payload records a passing test case in tests/fm-supervision-host.test.sh, but does not establish that it was driven against the live product. Provide a running product and drive the pass-t…
When ShellCheck hits a memory failure, lint retries without external sources and preserves genuine fallback findings. ⏸️ untested no The prior payload only records the memory-failure fallback test in tests/fm-lint.test.sh, not a live product lint run. Provide a running product and drive the ShellCheck memory-failure fallback scen…
  • bin/fm-test-run.sh tests/fm-pi-watch-extension.test.sh tests/fm-pi-branch-extension.test.sh tests/fm-calm-pi-extension.test.sh tests/fm-composer-lib.test.sh tests/fm-remote-job.test.sh tests/fm-watch-arm.test.sh tests/fm-lint.test.sh
  • FM_PI_BRANCH_LIVE_E2E=1 bin/fm-test-run.sh tests/fm-pi-branch-live-e2e.test.sh
  • bin/fm-test-run.sh tests/fm-supervision-host.test.sh (timed out at 1,200 seconds after the next-park watcher-reclaim case passed)
  • pi --version and installed Pi SDK version check (both 0.87.1)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

karotkriss and others added 24 commits September 30, 2026 07:19
…nguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…note (kunchenguid#6140)

* fix(bin): record Gerrit change URLs as close notes

Teardown's backlog_done_args hands every ship's recorded pr= URL to
fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a
canonical GitHub or Forgejo pull request. A Gerrit change URL therefore
left the item In flight after cleanup, and the pending backlog-close
record replayed into the same refusal at every session start.

fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads
as a Gerrit change into --note "Gerrit change <url>". The mapping sits
at the tasks-axi call rather than in the pending-close record, so
records already written with --pr replay to a close unchanged. The
captain-held retain path records the URL in its deliverable line and
skips the update --pr it cannot make.

* no-mistakes(review): Note retained Gerrit change URL when captain answers early

* no-mistakes(document): Document Gerrit change URL handling in captain-hold retention
* perf(remote): separate active job sampling from dispatcher cadence

* no-mistakes(document): Link remote wait timing to its authoritative contract

* no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed

* perf(supervision): reduce park, delta and dispatcher polling

* no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers
…henguid#6221)

* fix(bin): load backend sibling libraries under zsh

fm_backend_source kept each backend's sibling list in one space-separated
string and iterated it unquoted. zsh does not word-split an unquoted
expansion, so the readability check saw the whole list as one path and
refused every backend with more than one sibling. Hold the list in the
function's positional parameters instead, which needs no word splitting
in Bash 3.2, Bash 5, or zsh.

The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh
is installed.

* test: run the Calm mod suite on stock Bash 3.2

The suite injected shell values into its generated Node scripts with the
${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2
reports a bad substitution, so every case failed before it asserted
anything. Build each JavaScript string literal with JSON.stringify
through a small helper instead, which works on any Bash and is a valid
literal for any value.

* no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source

* test: narrow the zsh backend claim to name matching

Under zsh the adapters locate their siblings through BASH_SOURCE, so a
successful fm_backend_source is not a full load. Assert only what the
contract states, and pass js_string values after -- so node never reads
a leading-dash value as its own option.

---------

Co-authored-by: Nova Agent B <novaagentb@gmail.com>
…tatus scans (kunchenguid#5263)

* fix(bin): exclude a remote mate's own parent channel from self-home scans

A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home.
fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans.
The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path.
The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched.

* no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan

* no-mistakes(document): Document remote mate parent-channel scan exclusion

* ci: retrigger portable serial 4

* no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change

* no-mistakes(document): drop stale file-set copy in daemon catch-all comment
…6307)

* fix(bin): name the accepted verdict actors in fm-contributions help and refusal

* fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged
…chenguid#6306)

* fix(bin): recognise a clone root git names with different path spelling

fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone
root that git recorded with different casing (case-insensitive volume) was
skipped as not a clone root and never refreshed. Compare filesystem identity
instead, which also covers symlink spelling.

* fix(document): Remove stale clone-root comparison comment
)

* test(calm): pin Pi's regular TUI mode where pane assertions read scrollback

Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose
scrollable transcript is application-owned, so rows that leave the viewport
never enter terminal scrollback and tmux capture-pane -S can no longer see
them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag
exists so the transcript assertions keep reading real scrollback on both the
Pi 1.0.0 line and earlier Pi lines, which have no such flag and render
regular-only anyway.

* no-mistakes(document): Correct Pi TUI documentation and scrollback rationale
…ries (kunchenguid#6331)

* fix(bin): encode captain-hold reasons and reject self-inventory in complete

hold now stores a reason with parentheses, line breaks, or percent signs
through a reversible percent encoding that every reader decodes, instead of
refusing it. hold --origin records the origin on the held task, and complete
refuses the origin as its own inventory entry and an entry held for a
different origin; holds with no recorded origin are accepted and flagged.

* fix(review): Decode marked hold reasons consistently across readers

* fix(review): Remove unnecessary lifecycle test dispatch

* fix(review): Correct hold origin identity and inventory recovery

* fix(review): Record origins before placing backend holds

* fix(document): Clarify captain-hold validation and reason reader documentation

* fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed
* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits
…cessor already closed (kunchenguid#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check
* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check
…ommands (kunchenguid#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes kunchenguid#5601
Fixes kunchenguid#5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr
…#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged
…--external-sources (kunchenguid#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit
…tension log opt-in (kunchenguid#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.
…es (kunchenguid#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes kunchenguid#4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies
Upstream port point: 1f3e769

Brings in 20 upstream commits, including Pi 1.0 calm transcript
compatibility, Pi watcher continuity across successor gaps, titled Claude
top rules in the composer classifier, Pi trust approval for seeded
secondmate homes, remote-job polling churn, watcher-arm reclaim and the
ShellCheck memory-ceiling retry.

Conflicts in 10 files were resolved hunk by hunk, keeping both sides'
behavior; tests/fm-remote-job.test.sh gains upstream's SC2031 directive at
the fork's two later PATH-prefixed race commands, which the merge alone
made fire.
…y now reaches the arm’s terminal confirmation path, and the Pi HTML-renderer test supplies the Pi 1.0 renderer resolver. Updated the superseded-generation assertion to match its intended behavior. Relevant tests passed, including isolated watcher-lock rerun; a combined run had a transient watcher-lock failure under concurrency
…ke-queue file now counts as empty when checking an already-acknowledged generation. Updated the test to expect the established terminal-success exit status 4. `tests/fm-watch-arm.test.sh` passed; `git diff --check` passed
…tch by waiting for the stable “Reloading” prefix, while retaining the post-reload transcript and geometry assertions. The focused suite passed; the geometry E2E was skipped because Pi and tmux are unavailable here. `git diff --check` passed
@yelenplays
yelenplays merged commit 5c846aa into main Oct 3, 2026
38 of 39 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants